Modern Pathology
○ Elsevier BV
Preprints posted in the last 90 days, ranked by how well they match Modern Pathology's content profile, based on 22 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.
Zhu, M.; Li, A.; Safa, I.; Galera, P.; Hazoglou, M.; Vanderbilt, C.; Kamali, A.; Goldgof, G.; Veeraraghavan, H.; Jiang, J.; Ardon, O.; Geneslaw, L.; Dogan, A.
Show abstract
Pathologic diagnoses of hematopoietic diseases require immunohistochemistry (IHC) stains selected by pathologists upon preview of H&E-stained slides. This multi-step workflow can delay diagnostic turnaround time by days. Hence, we developed the Hematopathology Automatic Triaging System (HATS), which automates IHC panel ordering directly from H&E whole-slide images using pretrained pathology foundation model representations combined with attention-based multiple-instance learning. After the most comprehensive evaluation of pathology foundation models for hematologic malignancy classification to date, encompassing seven publicly available models, we trained HATS on 4,996 whole-slide images from 1,607 patients spanning the ten most common lymphoma diagnostic categories. HATS achieves 84% case-level subtype classification accuracy (0.962 ROC-AUC), translating to 92% IHC panel ordering accuracy. In a blinded reader study, HATS outperforms practicing pathologists at predicting lymphoma subtypes from morphology alone (85% vs 65%). In an independent real-world validation of 230 clinical cases, after directing 7 cases with scant tissue for manual review, HATS-ordered IHC panels were sufficient for diagnosis in 72.6% of cases. By automating the triaging step while preserving full pathologist oversight, HATS offers a safe and practical entry point for clinical AI adoption in pathology.
Ebbert, J. L.; Szymanski, J.; Perry, A.; Della Corte, D.
Show abstract
Automated Gleason grading now matches expert pathologists on the cohorts where systems are developed and tuned, but deployment-relevant gaps remain: whether an automated grade, applied without site-specific tuning or pathologist oversight, stratifies outcome comparably to expert grading on slides from unseen institutions and in cross-specimen applications. We tested this for disease-free interval (DFI), a curated recurrence endpoint. A production gland-level prostate diagnostic (PathTools Prostate v11.0) was applied frozen and uncalibrated to 298 diagnostic whole-slide images from 274 TCGA-PRAD radical-prostatectomy patients, a cohort outside its development distribution and needle-core-biopsy training data, contributed by 25 source sites under heterogeneous digitization; tissue was detected automatically with no expert region annotation. From the output we derived an ISUP grade group and continuous high-grade content, and evaluated each grade as a standalone predictor of DFI (24 events) by Harrell's c-index with 95% bootstrap confidence intervals, a paired between-method bootstrap, and Kaplan-Meier curves with the log-rank test. The automated grade reproduced the clinical grade group at quadratic-weighted kappa = 0.62 (95% CI 0.53-0.70; 48% exact, 86% within one group), within the expert inter-observer range. As the sole predictor it stratified recurrence (log-rank p = 0.022; c-index 0.69, 95% CI 0.58-0.79), and the continuous high-grade fraction was robustly prognostic (hazard ratio 1.37 per SD, p = 0.029; c-index 0.71, 0.61-0.81). Standalone discrimination was not statistically separable from the clinical grade (c-index 0.78, 0.69-0.86; paired {triangleup} c-index spanning zero), and in a joint model the automated grade added nothing beyond it, consistent with both measuring a shared morphological axis. From a single out-of-distribution slide with no pathologist oversight, the automated grade provides standalone recurrence stratification not statistically separable from whole-gland expert grading, demonstrating robust generalizability beyond training data; reported as a continuous high-grade fraction, it offers reproducible, expert-free, grade-equivalent risk stratification for harmonizing large archival or genomically-profiled cohorts.
Calapaqui Teran, A. K.; Gonzalez Bernad, A. A.; Cobo Cano, M.; Sanchez Magdaleno, L.; Marcos Gonzalez, S.; Delgado Bolton, R. C.; Moustafa Calvo, J.; Gomez Roman, J. J.; Lara, L.
Show abstract
We present PRECISE (PRostate Expert-annotated Contiguous IHC-H\&E Serial sEctions), a hybrid histopathology dataset of paired hematoxylin and eosin (H\&E) and immunohistochemistry (IHC) whole-slide images (WSIs), comprising 37 prostate core needle biopsies from 25 patients, each with matched H\&E and CKAPM+racemase staining. To the best of our knowledge, this is the first publicly available dataset offering spatially harmonized, pixel-level expert annotations across both staining modalities in prostate biopsy WSIs - directly mirroring the two-stage (H\&E-then-IHC) clinical diagnostic workflow used to resolve morphological uncertainty, restricted to cases in which that workflow reached diagnostic consensus. The dataset contains 24,387 annotations spanning seven diagnostically critical classes: malignant glands, benign glands, stromal tissue, intraductal carcinoma (IDC-P), high-grade prostatic intraepithelial neoplasia (HGPIN), atypical intraductal proliferation (AIP), and tissue artifacts. Unlike existing resources, which focus on binary tumor classification or lack IHC pairing, this dataset captures the full morphological spectrum encountered in routine prostate pathology, including rare precursor lesions and confounding entities underrepresented in current benchmarks. Annotations were validated through a structured three-stage consensus by two expert uropathologists, with IHC serving as biological ground truth for boundary definition. PRECISE is designed as a robust benchmark for multimodal semantic segmentation and self-supervised learning, and is openly released to promote reproducible research and accelerate AI-assisted diagnosis in prostate cancer.
Ebbert, J. L.; Perry, A.; Szymanski, J.; Della Corte, D.
Show abstract
Background: Deep-learning systems for Gleason grading are developed almost entirely on high-end clinical scanners and on cohorts from a small number of Western institutions, yet deployment increasingly involves other devices and other populations. These two distribution shifts, device and population, are rarely tested together on the same physical slides. The PAR dataset, from Erbil, Iraq, digitizes each biopsy on three scanners and provides three distinct pathologist grades, so it permits both tests at once on a Middle Eastern cohort. A concurrent study by the dataset originators validated a task-specific model and two foundation models on PAR; we complement it by testing an inde-pendently developed detect-then-grade pipeline and by separating scanner effects on detection from scanner effects on grading. Methods: We applied one fixed model de-veloped on North American and European material to all 1017 whole-slide images (339 slides from 185 patients, three scanners; 49.6% clinically significant cancer) with no scanner-specific or population-specific tuning. We measured cancer detection (area under the ROC curve of the predicted cancer-tissue fraction), all-slide ISUP agreement of the deployed detect-then-grade pipeline (quadratic-weighted kappa, QWK), and grading agreement on pathologist-confirmed cancers, at the slide level and, because a case carries up to two slides, at the patient level. The reference reader was S.A.; thresholds and operating points were cross-validated leave-one-out; scanners were compared by paired within-biopsy bootstrap and confidence intervals confirmed by patient-cluster bootstrap. Results: Detection was statistically equivalent across scanners (AUC 0.987 to 0.991; paired differences at most 0.003) and transferred to this non-Western cohort with no per-population tuning. At a 95% sensitivity operating point the deployed pipeline reached cross-validated all-slide QWK of 0.86, 0.81, and 0.86 (Grundium, Hamamatsu, Leica), matching the inter-pathologist ceiling of 0.81, against 0.23 to 0.62 for the ungated model. Grading of confirmed cancers was scanner dependent: the compact Grundium (0.63) did not differ from the clinical Leica (0.67; paired difference 0.04, 95% CI -0.03 to 0.11), while both exceeded Hamamatsu (0.44). Results held at the patient level, with grading somewhat lower for two scanners; the two slides of a case disagreed in grade in 43% of cases, and patient clustering did not widen the intervals. Conclusions: Can-cer-tissue fraction is a triage signal robust across scanner and transferable to an un-derrepresented population for detection, while grading is the scanner-sensitive step. Prostate grading models should be deployed as a detect-then-grade pipeline, with grading validated per device and confirmed on the local population.
Shah, N. A.; Sarwar, M.; Ullah, E.
Show abstract
Background: Homologous recombination deficiency (HRD) is clinically imperative in high-grade serous ovarian carcinoma (HGSOC), particularly because of its association with platinum sensitivity and benefit from poly(ADP-ribose) polymerase inhibitor (PARPi) therapy. However, public datasets rarely contain a complete combination of diagnostic haematoxylin and eosin (H&E) whole-slide images (WSIs), validated clinical HRD assay results, genomic scar scores, BRCA1 promoter methylation data, and treatment-response outcomes. This creates a major barrier for computational pathology studies seeking to develop clinically interpretable models of HRD or PARPi response from routine histology. Objective: We performed an exploratory, leakage-controlled computational pathology benchmarking study to evaluate whether H&E WSIs from TCGA-OV contain a measurable morphology-linked signal associated with research-grade molecular HRD labels, and whether label refinement and pathology foundation-model embeddings alter predictive performance. Methods: We assembled a frozen-primary TCGA-OV WSI cohort comprising 717 tissue-section/biospecimen slides from 316 patients. Diagnostic FFPE DX slides were excluded from model selection because of complete patient overlap with the frozen-primary cohort. Two HRD labels were evaluated: an initial mutation-only molecular label based on BRCA/HR-gene mutation evidence, and a refined methylation-enhanced molecular label that additionally incorporated BRCA1 promoter methylation. Feature extraction was performed using ResNet50, UNI, CONCH, Virchow2, Phikon-v2, and UNI2-h encoders. Patient-level attention-based multiple instance learning (ABMIL) was used with patient-as-bag modelling. Evaluation used patient-level grouped 5-fold x 5-repeat stratified cross-validation, with 25 folds total, bootstrap confidence intervals, and patient-level leakage control. Results: The initial mutation-only label classified 78 patients as positive and 238 as negative. The refined methylation-enhanced label recovered 33 additional positives, resulting in 111 positive and 205 negative patients. Patient-level ABMIL using UNI2-h features achieved the strongest performance for the refined label, with AUROC 0.634 (95% CI 0.571-0.698), AUPRC 0.468 (95% CI 0.390-0.562), balanced accuracy 0.597, sensitivity 0.532, specificity 0.663, F1 score 0.494, and Brier score 0.233. The calibrated threshold was 0.512, yielding TN=136, FP=69, FN=52, and TP=59. Comparative models showed lower discrimination, including UNI2-h with the initial label (AUROC 0.628), Phikon-v2 refined (0.582), Virchow2 refined (0.582), CONCH initial (0.587), ResNet50 refined (0.570), and clinical baselines (AUROC 0.54-0.57). Conclusions: TCGA-OV H&E WSIs contain a modest but reproducible morphology-linked signal associated with research-grade molecular HRD status. However, the AUROC around 0.63, absence of clinical HRD assay labels, lack of genomic scar endpoints in the implemented workflow, and absence of PARPi/platinum response targets prevent clinical interpretation. This study should be interpreted as a proof-of-concept benchmarking framework and methodological foundation for future H&E-based predictive modelling in clinically curated PARPi response cohorts.
Doeleman, T.; Brussee, S.; Valkema, P.; Kempf, W.; Vermeer, M.; Kers, J.; Wynaendts, L.; Kerckhoffs, K.; de Jonge, M.; Nguyen, A.; Peters, E.; Wobser, M.; Rauert-Wunderlich, H.; Rosenwald, A.; Stadler, R.; Jansen, P.; Battistella, M.; Roccuzzo, G.; Quaglino, P.; Schrader, A.
Show abstract
Background Histological diagnosis of early-stage mycosis fungoides (MF) is hindered by profound overlap with benign inflammatory dermatoses (BIDs), leading to diagnostic delays and extensive ancillary testing. We developed MIMIC (Multiple Instance-learning for Identification of Mycosis fungoides In Cutaneous biopsies), a weakly supervised deep learning model designed as a triage tool at initial H&E whole slide image (WSI) review to distinguish classic patch and plaque stage MF from BIDs. We externally validated the model and evaluated its clinical utility. Methods In this retrospective multicentre study, we trained a base model using weakly supervised attention based multiple instance learning on 3,339 WSIs from two Dutch centres. Crucially, all MF training labels were derived from a deeply phenotyped national cohort featuring strict multidisciplinary expert panel consensus diagnoses (the clinical gold standard). Transportability was evaluated on 371 WSIs from four independent European centres. A blinded reader study on 171 WSIs compared morphology only performance of MIMIC with 11 (dermato-)pathologists. We then retrained an updated model on all retrospective multicentre data and assessed clinical utility in a strictly held out, consecutive Utrecht cohort (2022-2023; 486 accessions, 863 WSIs). Primary analysis focused on classic MF versus BIDs (453 accessions). Decision curve analysis, using Platt scaled probabilities to correct for spectrum bias, evaluated net benefit at a prespecified, safety oriented threshold of 0.04. Findings The base model showed good multicentre transportability (mean centre specific AUROC 0.91; pooled AUROC 0.84). In the reader study, MIMIC achieved an AUROC of 0.87, exceeding the mean pathologist AUROC (0.79) and the best individual reader (0.83). In the consecutive MF versus BID cohort, the updated model achieved an AUROC of 0.87 (95% CI 0.81-0.92). At the 0.04 threshold, sensitivity was 97.8% (44/45 MF cases) and specificity 50.2%, reducing unnecessary ancillary workups by 39.9 per 100 screening cases versus a test all strategy. Interpretation By identifying nearly half of BIDs as low risk while preserving near complete sensitivity for classic early stage MF in a European digital pathology workflow, this unimodal H&E approach offers a scalable digital solution to reduce defensive ancillary testing and accelerate the diagnostic journey for patients with MF. Further validation is needed in non European centres and in populations with darker skin phototypes.
Kambou Kountchou, K. D. K. K.; Tommo Tchouaket, M. C.; Moko Fotso, L. G.; Fokou Bomgning, B. N.; Fippo Fitime, L.; Talom Teumadjou, A.; Routoube, M.; Efakika Gabisa, J.; Ngoufack Jagni Semengue, E.; Nka, A. D.; Kae, A. C.; Dobgima Pisoh, W.; Deutou, L.; Takou, D.; Fainguem, N.; Sosso, S. M.; Kamgaing Simo, R.; Yagai, B.; Tabola Fossa, L.; Perno, C.-F.; Colizzi, V.; Enow-Orock, G.; Fokam, J.; Terrinoni, A.; Kuiate, J.-R.
Show abstract
Background: In resource-limited settings, a critical bottleneck in cervical cancer prevention is the lack of practical strategies to triage high-risk human papillomavirus (HR-HPV)- positive women. Therefore, this study aimed to develop and internally validate a genotype-specific risk stratification model. Methods: A cross-sectional study enrolled 555 women in Cameroon. Data collection integrated cervical cytology and HPV genotyping using Abbott m2000rt and Sacace multiplex systems. An iterative modeling approach with bootstrap validation was used to develop the model and address model instability. HR-HPV genotypes were transformed into a hierarchical risk variable due to sparsity and integrated with significant predictors. The final model was translated into a scoring system, and the risk gradients and performances were evaluated at two thresholds. Data was analyzed using SPSS 27.0. Results: The mean age was 44.8 years, and the prevalence of HR-HPV was 26.5% (147/555). The final model, incorporating HPV categories, age, and tobacco, demonstrated moderate discriminative ability (AUC=0.702, 0.642-0.762) with a good calibration (Hosmer-Lemeshow {chi}{superscript 2}=4.05, p=0.399). The scoring system assigned women to risk groups based on their total scores which produced a clear monotonic risk gradient; the observed probability of high-grade lesions/cancer ranged from 15% (score 0) to >65% (score [≥]4). At a conservative threshold ([≥]4 points), 4.7% (26/555) of women were classified as high-risk, concentrating 46% (6/13) of cancers (positive predictive value[PPV]=58%) while a sensitive threshold ([≥]3 points) had 16.8% (93/555) high-risk, concentrating 77% (10/13) cancers (PPV=38%). Both thresholds maintained a high negative predictive value (>95%). Conclusion: This bootstrap-validated, risk-stratification tool is a proof-of-concept in resource limited settings that assigns HR-HPV-positive women to distinct management pathways using three variables. After refining through a longitudinal study and external validation, this scoring system can improve the efficiency of cervical cancer screening programs in low-resource settings.
Hofstraat-Boersma, R.; du Long, R.; Buzzanca, G.; Abiola, A. A.; Albadri, S.; Ali, Z.; Altaleb, A.; Angioi, A.; Banu, S. G.; Barry, M.; Bhalodia, A. R.; Bianco, P.; Broecker, V.; Buelow, R.; Chauveau, B.; Chen, G.; Cheunsuchon, B.; Crisi, G. M.; Daneshvar, S.; Dendooven, A.; Dokouhaki, P.; Drachenberg, C. B.; Farris, A. B.; Ferlicot, S.; Florquin, S.; Fontana, F.; Gibier, J.-B.; Gibson, I. W.; Gujarathi, S.; Hendricks, A. R.; Husain, S.; Islam, J.; Ismail, W.; Jagannathan, G.; Klager, J.; Kozakowski, N.; Krizova, A.; Kurien, A. A.; Kwon, B.; L'Imperio, V.; Ledesma, F. L.; Low, J. P.; Martin, J
Show abstract
Background Diagnostic interpretation of kidney allograft biopsies using the Banff classification remains variable, but the determinants of this variability are not fully defined. We performed a global, fully digital multi-reader study to identify the principal drivers of disagreement in Banff-based assessment. Methods Thirty six kidney transplant biopsies were independently scored by 67 renal pathologists on a standardized digital platform. Readers assessed Banff lesions on hematoxylin and eosin, periodic acid Schiff, and Jones' silver stains; final diagnostic categories were assigned using prespecified Banff-based decision rules. Interobserver agreement was quantified with Gwet's agreement coefficient (AC) statistics. Determinants of diagnostic agreement were evaluated) using pairwise mixed-effects logistic regression, and reader similarity was examined by principal component analysis (PCA) with post hoc molecular annotation. Results Agreement for final diagnostic categories was moderate (Gwet's AC1, 0.55; 95% CI, 0.47 - 0.63). Lesion-level agreement varied substantially, with lowest agreement for selected threshold-dependent inflammatory or semi-quantitative lesions, including interstitial inflammation in areas of IFTA, peritubular capillaritis and arteriolar hyalinosis. Diagnostic concordance differed markedly across biopsies, indicating strong case-level heterogeneity. In pairwise models, differences in active inflammatory and vascular lesion scoring were the strongest correlates of diagnostic disagreement; reader experience and geography contributed minimally. Principal component analysis showed reader variation was organized along two dominant axes: a rejection-calling threshold axis linked mainly to tubulointerstitial inflammatory injury, and a T cell-mediated (TCMR/TI) and antibody-mediated/microvascular (AMR/MVI) inflammation-oriented phenotypic classification axis. Conclusion Interobserver variation in Banff-based kidney transplant biopsy assessment is structured rather than random and driven mainly by how readers threshold and integrate key inflammatory lesion compartments rather than experience or geographic location.
Paulikat, M.; Bosch, C.; Aswolinskiy, W.; Caixeta Borges, I.; Nauschuette, L.; Aichmueller, C.; Schmidt, D.; Bussmann, H.; Kalteis, S.; Zapukhlyak, M.; von Knebel Doeberitz, M.; Kloor, M.
Show abstract
Accurate grading of cervical biopsies on Hematoxylin and Eosin (H&E) stained whole slide images (WSIs) is essential for distinguishing high grade lesions from low grade changes, yet this process is subject to considerable inter-observer variability. In this study, we evaluate a foundation model-based multiple instance learning (MIL) pipeline for binary high-grade squamous intraepithelial lesion (HSIL) detection on H&E stained WSIs. We benchmark our Athena foundation model against four state-of-the-art pathology foundation models: H-optimus-0, Hibou-L, Midnight-12k and Virchow, across datasets from five different countries: Portugal, Cambodia, Germany, Poland and Scotland. Athena achieved the highest mean area under the curve (AUC) (0.931) with the lowest cross-country variability (STD = 0.022). Furthermore, we compared the model's diagnostic performance to that of trained pathologists on a dataset with p16-confirmed ground truth. Our model improved sensitivity from 84% to 95% while maintaining comparable specificity (85% vs. 84%). Failure analysis revealed that the model's errors were concentrated at the diagnostic boundary between low-grade and high-grade lesions, whereas pathologists' errors spanned a broader range of misclassifications. These findings show the potential of foundation models for cervical cancer screenings worldwide.
Kang, Y.-J.; Jun, S.-Y.; Kim, S.
Show abstract
Background. Breast cancer treatment depends on histopathological features, such as grade and receptor-defined subtype; however, specialist pathologist access is constrained when the workforce is limited. Commercial multimodal large language models (MLLMs) accept hematoxylin and eosin (H&E) image tiles through paid interfaces without local hardware or fine-tuning. However, prior pathology evaluations addressed only coarse tasks. Whether they reach treatment-determining accuracy and whether vendors agree remain unclear. Methods. We aimed to evaluate three vendor-designated flagship MLLMs (Claude Sonnet 4.6, Gemini 2.5 Pro, GPT-5.5) in 427 invasive breast cancer cases. Each case went to all three with identical H&E tiles and prompts, and the subtype was inferred in the second call. The reference was an institutional sign-out report of an immunohistochemistry-derived subtype. We calculated the concordance, sensitivity, specificity, Cohen's kappa, and pairwise McNemar and Bowker tests. Findings. Claude ranked highest by raw histologic-type concordance but lowest by kappa, classifying all 23 lobular and seven micropapillary carcinomas as invasive breast carcinoma of no special type. The models anchored the Nottingham grade to three modal grades. None of the models reliably identified human epidermal growth factor receptor 2-positive disease. The failure direction was vendor-specific: Claude and GPT-5.5 were under-detected, whereas Gemini was over-called. Twelve prompt variants (4,056 calls) did not recover sensitivity. Interpretation. No current commercial MLLM reaches deployment-ready accuracy for any treatment-determining feature of breast pathology. As each vendor fails in its own fixed direction, changing vendors alters the type of error rather than removing it; therefore, the value of these models is assistive rather than autonomous. At USD 0.20-0.50 per case, they may serve as supervised draft generators that leave the diagnosis with the pathologist.
Wang, E.; Grenier, K.; Savadjiev, P.; Poenaru, D. D.
Show abstract
Background. Definitive diagnosis of Hirschsprung's disease (HD) requires pathological identification of enteric ganglion cells. This process is time-consuming and subject to inter-observer variability. Artificial intelligence (AI) tools have the potential to standardize and accelerate this workflow, but no study has determined which AI approach best serves intraoperative HD pathology diagnostics. Method. This study compared the U-Net and You Only Look Once version 26 (YOLO26) frameworks for ganglion cell detection using a single-centre retrospective dataset of 54 whole-slide images (WSIs) from rectal biopsies. WSIs were tiled into 397,731 image patches (128x128 pixels), further partitioned into training (70%), validation (15%), and testing (15%) sets. Models were evaluated on tile- and patient-level diagnostic metrics and processing latency. Results. The U-Net achieved a tile-level sensitivity of 82.9%, showing no statistically significant difference compared to YOLO26 (79.1%; p = 0.097). However, YOLO26 demonstrated a statistically significant advantage in tile-level specificity (96.1% vs. 93.9%; p < 0.001) and reduced mean inference latency (7.64 ms vs. 11.57 ms/tile). At the patient level, both models achieved 100% diagnostic sensitivity. Despite low patient-level specificity (0.0% U-Net; 11.8% YOLO26), the tissue-level diagnostic burden of false positives was 6.00% for U-Net and 3.50% for YOLO26. Conclusion. The U-Net is preferred when nominal gains in sensitivity are prioritized, while the YOLO26 is an alternative that optimizes efficiency and false positive suppression. Both models serve as robust screening filters to augment the pathologist's workflow and should be selected based on workflow requirements. Prospective validation on larger, multi-centre datasets is required before clinical implementation.
Fuller, T. D.; Polidoro, R. B.; Strand, D. W.; Arrizabalaga, G.; Jerde, T.
Show abstract
Background: Chronic inflammation is the most common histological feature in Benign Prostatic Hyperplasia (BPH), and T cells are a key component of immune infiltrate. Advanced BPH is commonly associated with the formation of nodules, but it remains unclear whether a link exists among T cell infiltration, nodular development, and BPH progression. Using a Toxoplasma gondii (T. gondii) model and human specimens, we characterize the subtypes of T cells present during prostatic hyperplasia and their association with nodular development of the prostate. Methods: Male CBA/j mice were intraperitoneally infected with T. gondii parasites, and flow cytometry was performed on the prostate to quantify the number of CD4+ and CD8+ T cells. Histology was used to score microglandular hyperplasia (MGH), and immunofluorescence was used to quantify and examine the locality of CD4+ and CD8+ T cells and compared that to human BPH tissue. Results: We found that infecting male mice with T. gondii resulted in an increase of both CD4+ and CD8+ T cells in the prostate acutely and that CD8+ cells remained sustained at chronically. We also established the presence of glandular nodule formation at this timepoint through hematoxylin and eosin (H&E) staining. Immunofluorescence revealed that CD8+ cells were found proximal to forming glandular nodules relative to non-nodular glands. We also found more CD8+ cells localized to non-nodular glands in nodular BPH tissue versus non-nodular BPH tissue. Finally, we discovered a higher prevalence of CD8+ cells in T. gondii IgG+ patients than in IgG- patients. All T. gondii IgG+ patients exhibited nodular BPH, whereas all but one IgG- patient exhibited non-nodular BPH. Conclusions: This study is the first to investigate the presence and location of CD4+ and CD8+ T cells within nodular and non-nodular BPH glands. We found an association of the presence of CD8+ T cells with nodular progression. This association held true in human prostate tissue. Translationally, CD8+ T cells may enhance nodular BPH progression, and T. gondii infection may promote this CD8+ T cell-mediated response.
Buzzanca, G.; Pala, C.; He, J.; Hofstraat-Boersma, R.; Tammaro, A.; van Midden, D.; Buelow, R.; Hoelscher, D. L.; Muehlfeld, A. S.; Koeller, m.; Kozakowski, N.; Boehmig, G.; Halloran, P. F.; van der Helm, D.; Meziyerh, S.; Venhuizen, J.-H.; Haitjema, S.; Dijkstra, J.; Hilbrands, L. B.; Steenbergen, E. J.; van Zuilen, A. D.; Nurmohamed, A. S.; Bemelman, F. J.; Bruns, I. B.; Callegaro, G.; van de Water, B.; Pieters, T. T.; Breimer, G. E.; Rossi, G. M.; Fiaccadori, E.; Maggiore, U.; Roelofs, J. J. T. H.; Testa, F.; Fontana, F.; Abiola, A. A.; Delsante, M.; Corthals, G. L.; Peters-Sengers, H.; Ngu
Show abstract
Accurate, reproducible interpretation of kidney allograft biopsies is critical for diagnosis of graft injury to guide prognosis and management. The international Banff classification is a consensus diagnostic system based on semiquantitative histological lesion scoring on either extent or severity of kidney transplant biopsies. However, pathologist scoring is limited by substantial interobserver variability, constrained scalability, and the inherent nature of the scoring system itself. Here we present BanffNET, a weakly supervised, probabilistic deep learning framework that combines self-supervised feature extraction with a novel Bayesian multiple-instance learning framework to predict (continuously) the full spectrum of Banff lesion scores directly from whole-slide images (WSIs). Using lesion-specific aggregation functions tailored to localized (modeling lesion severity) and diffuse pathologies (modeling lesion extent), BanffNET generates interpretable, patch-level probability maps and calibrated slide-level scores. BanffNET's performance was assessed relative to consensus, biological correlates of rejection and clinical outcome, demonstrating superior consistency, transportability and generalization. Trained on 7,249 WSIs from three cohorts, BanffNET demonstrates consistent performance on 11,028 WSIs across five external test sets, performing on par or exceeding expert consensus across lesions. BanffNET scores align more closely than pathologist Banff scores with molecular profiles of rejection, offering a transparent, biologically grounded framework for computational pathology with relevance beyond transplantation.
Liu, Y.; Loneman, D.; Bready, B.; Nemirovsky, D.; Cohen, A.; Wang, X.; Stein, E.; Zhang, Y.; Derkach, A.; Hasserjian, R. P.; Xiao, W.
Show abstract
The 5th Edition of the World Health Organization Classification of Haematolymphoid Neoplasms (WHO5th) and the 2022 International Consensus Classification (ICC) both recognize myelodysplasia-related acute myeloid leukemia (AML-MR) as a diagnostic entity increasingly defined by integrated genomic data. Although largely concordant, the two classifications differ in various ways that should be resolved to achieve future harmonization. To address the areas of uncertainty, we retrospectively analyzed 615 newly diagnosed AML cases from adult patients treated at two large cancer centers. We demonstrate that AML-MR, whether defined by gene mutations (MR-GM) or cytogenetic abnormalities (MR-CGA), constitutes a prognostically distinct group with inferior outcome compared to most AML subtypes, second only to TP53-mutated or EVI1-rearranged AML. Isolated RUNX1 mutations were not associated with antecedent myeloid neoplasia. Neither the number of mutated MR genes nor their variant allele frequency independently impacted outcomes. Trisomy 8 and del(20q) did not confer inferior outcomes and may warrant exclusion from MR-CGA. Complex karyotype without TP53 mutations did not worsen outcomes within AML-MR and may be considered equivalent to other MR-CGA. The adverse prognosis of AML-MR appeared to be at least partly driven by ASXL1 and/or EZH2 mutations. These findings provide evidence toward a unified schema across the WHO5th and ICC.
Wendt, J. R.; Adams, K. M.; Moreno, R.; Hossan, M. S.; Stram, A.; Lin, E. S.; Kersten, L.; Kratz, J. D.; Roy, M.; McGregor, S. M.; Lang, J. D.
Show abstract
Patient-derived organoids (PDOs) have transformed translational cancer research, allowing tractable models that better represent clinical features than traditional immortalized cell lines. Here we describe two PDOs with differential responses to carboplatin derived from sequential ascites fluid collections from a patient with high-grade mullerian carcinoma, that could not be further subclassified on the omental biopsy. Uterine origin was clinically excluded by pelvic imaging/CT scan of the uterus and absence of vaginal bleeding. Successful derivation from independent collections enabled comparison of intra-patient heterogeneity across sequential ascites samples and demonstrates that PDO efficiency rate is at least partly patient-specific or tumor-dependent. We performed long-read whole genome sequencing on the two PDOs, OC104 and OC109, to better characterize the structural variant landscape while also obtaining information on single nucleotide variants and DNA methylation. In addition to confirming single nucleotide variants noted in clinical sequencing (TP53, KRAS, SPOP, PPP2R1A, KMT2D), we identified additional variants in TSC2, NCOR2, and CTNNA2 that are predicted to be likely pathogenic. The spectrum of mutations, particularly the coincident KRAS and TP53, highlighted unexpected overlap with ovarian mucinous carcinoma. We also identified larger insertions and deletions that result in non-synonymous variants in MUC5AC, TPRX1, and BMX, as well as four translocation events, including two that could not have been resolved with short-read sequencing. Differentially methylated promoters between the two PDOs include 201 oncogenes and tumor suppressor genes, with HNF1A, MSI2, and SETBP1 having methylation directions consistent with these genes' roles in platinum response differences observed between the PDOs. Notably, the clonal nature of PDOs produced from two samples taken one week apart is important for the field to appreciate, particularly since they have clonal differences in platinum response. The temporal differences in clonality may indicate a limitation of low volume sampling, however may provide opportunity to longitudinally predict clinical outcomes. We also demonstrate the ability of long-read sequencing to add detail into the genomics and epigenetics of ovarian cancer.
Heilijgers, F.; Le, H. A.; Coudray, N.; Karimkhan, A.; Chen, D.; Peeters, K. C. M. J.; Hacking, S.; Mesker, W. E.; Tsirigos, A.; UNITED collaboration,
Show abstract
H&E whole-slide images capture prognostic information encoded in tumor morphology and the surrounding microenvironment, but these signals remain difficult to extract and interpret at scale. Here, we developed a self-supervised computational pathology framework to predict disease-free survival in colorectal cancer and link model-derived risk to interpretable histomorphology and spatial tumor biology. Using a multicenter developmental cohort spanning colorectal adenomas and invasive colorectal cancer, we trained HPL-PanColon, a self-supervised representation model, to extract tile-level embeddings and identify recurrent histomorphological phenotype clusters across the adenoma-carcinoma spectrum. Compared with general-purpose pathology foundation models, HPL-PanColon yielded representations with reduced institution- and dataset-specific batch effects. We then applied HPL-PanColon to a global survival cohort of 1,024 colorectal cancer patients in a leave-one-institution-out framework, using tile embeddings to train an attention-based survival model and derive the Colon Histomorphology Prognostic Score (CHiPS). CHiPS stratified patients by disease-free survival and provided complementary prognostic information to a UICC TNM-informed clinicopathological model, increasing the c-index from 0.683 to 0.706. Integrating model attention with phenotype assignments traced CHiPS-associated risk to pathologist-recognizable tissue patterns, with high-risk regions enriched for desmoplastic, stromal, and fibroinflammatory morphologies and low-risk regions reflecting tumor-rich epithelial glandular patterns. Spatial transcriptomic analysis further linked high-risk morphologies to fibroblastic, perivascular, myofibroblastic, and immune-reactive tumor microenvironment programs, while low-risk morphologies mapped to epithelial and tumor-enriched regions. These findings establish a scalable framework for interpretable histology-based prognosis and spatial biological discovery in colorectal cancer.
Connelly, J.; Hernando, B.; Luft, J.; Anderson, C. J.; Bankhead, P.; Connor, F.; Aitken, S.; Liver Cancer Evolution Consortium, ; Semple, C. A.; Flicek, P.; Odom, D. T.; Taylor, M. S.; Aitken, S. J.
Show abstract
Background & AimsHaematoxylin and eosin (H&E) staining remains the diagnostic gold standard for solid cancers, including hepatocellular carcinoma, and is increasingly complemented by genomic profiling for precision medicine. Inferring genomic alterations directly from H&E images could streamline testing, but heterogeneity and biases in human training data limit interpretation of genotype-phenotype associations. Here, we aimed to relate histologic to genomic pathology to provide biological explainability for mutation prediction models and assess the impact of germline variation on model performance. MethodsWe analysed 597 murine liver tumours with matched whole-genome sequencing and histopathology (163,835 image tiles; 22.9 million nuclei). Our controlled in vivo design accounted for germline variation, biological sex, and causal mutagen (N-diethylnitrosamine), removing confounding factors present in human cohorts. We trained and evaluated deep learning and supervised machine learning models to predict germline variation and cancer driver alterations from H&E. ResultsModelling accurately predicted germline and somatic alterations from histology, at both locus-specific and genome-wide scales. Quantitative image analysis revealed an unexpected association between Egfr driver mutations and hepatic steatosis, linking genotype to an interpretable morphological phenotype. While model performance declined when applied to tumours from unrepresented genetic backgrounds, this limitation was biologically informative, revealing strain-dependent differences in tumour evolution, notably the prevalence of whole-genome duplication. ConclusionsMachine learning integration of histological and genomic pathology enables accurate, interpretable inference of genetic alterations from H&E, potentially reducing reliance on costly ancillary molecular assays. Our predictions are supported by human-interpretable biological features, addressing concerns around "black-box" technologies. However, caution is required when applying such methods to samples with a genetic background that, even if closely related, is beyond the genetic horizon of training data.
Tahir, W.; Shamshoian, J.; Tauber, J.; Clinton, L. K.; Griffin, M.; Shah, C.; Singh, G.; Fahy, D.; Sucipto, K.; Brosnan-Cashman, J.; Altepeter, T. A.; Bhattacharya, S.; Crandall, W.; Duan, C.; Gale, J. D.; Gupta, V.; Haarmann, H.; Harpaz, N.; Hooper, A. T.; Horowitz, J.; Hurtado-Lorenzo, A.; Hussaini, B. E.; Jairath, V.; Jones, A.; Kostiuk, B.; Kukreja, A.; Laroux, F. S.; Lissoos, T.; McBride, R. B.; Najdawi, F.; Nayyar, A.; Osterman, M. T.; Panchal, P.; Ruane, D.; Travis, S.; Visvanathan, S.; Wilson, L.; Jayson, C.
Show abstract
In clinical trials for ulcerative colitis (UC), pathologists assess disease severity through standardized histological indices, including the Geboes Score, Robarts Histopathology Index (RHI), and Nancy Histologic Index (NHI). Despite strong associations with clinical outcomes, histologic scoring suffers from inter- and intra-reader variability, and consensus criteria for histologic remission remain uncertain. Through a consortium approach, we developed an artificial intelligence-based measurement (AIM) tool for scoring histology in UC mucosal biopsies (AIM-HI UC). This model, trained on a large dataset of UC biopsies (N=10,230), utilizes additive multiple instance learning models leveraging PLUTO, a pathology foundation model, that predict each of the Geboes subgrades, from which the Geboes grade-level score, RHI, and NHI can be calculated. Evaluation of this model on a standalone verification set including clinical trial specimens established algorithm non-inferiority and/or superiority relative to standard qualified pathologists through comparison of algorithm-consensus and pathologist-consensus agreement metrics (non-inferior if difference >-0.1, superior if difference >0, inclusive of confidence intervals). AIM-HI UC was determined to be non-inferior to pathologists (N=3) for the prediction of all seven Geboes subgrades, grade-level Geboes, RHI, NHI, histologic improvement (GS<3.1), 2A histologic remission (GS<2A.0), and 2B histologic remission (GS<2B.0). AIM-HI UC was superior to pathologists for several Geboes subgrades (GS 0, GS 1, GS 2B, and GS 5), as well as grade-level Geboes, RHI, and positive percent agreement of 2A histologic remission. The model was shown to be greater than 99% repeatable for all histologic scoring metrics examined. Model-derived scores were shown to strongly correlate with canonical histologic features of inflammation, including the proportion of total epithelium that is inflamed (Spearman r=0.83; p<0.01), the proportion of neutrophils localized within crypt epithelium (Spearman r=0.83, p<0.01), and the amount of mucosal area classified as erosion or ulceration (Spearman r=0.80, p<0.01). Overall, these results suggest that AIM-HI UC has the potential to improve consistency of UC histology interpretation, providing a path toward standardization of UC histology scoring in clinical trials.
Yan, R.; Gao, G.; Song, A. H.; Hsieh, H.-C.; Zhao, Y.; Almagro-Perez, C.; Brenes, D.; Chow, S. S. L.; Shen, J.; Reddi, D. M.; True, L. D.; Lal, P.; Madabhushi, A.; Mahmood, F.; Liu, J. T. C.
Show abstract
Non-destructive 3D pathology enables high-resolution slide-free imaging of intact clinical specimens, providing comprehensive visualization of tissue structures beyond what conventional slide-based 2D histopathology can provide. However, the scale and complexity of volumetric datasets make exhaustive manual review impractical, motivating AI-assisted triage methods to select a small number of high-risk 2D slices for pathologist review. While prior triage models have shown promise, interpretability is poor and performance can be suboptimal, especially in the nascent field of 3D pathology in which labeled data is limited. We present SCOPE, a Segmentation-guided CrOss-slice PrototypE learning framework for comprehensive risk assessment of 2D levels within 3D pathology datasets. SCOPE combines (i) clustering-based pretraining on large-scale unlabeled volumetric data to initialize morphology-aware prototypes, (ii) segmentation-derived structural priors from publicly available models to guide proto-type learning, and (iii) cross-slice (2.5D) prototype aggregation across neighboring slices to generate slice-level risk predictions. In prostate and esophageal data cohorts, SCOPE consistently outperforms attention-based and prototype-based multiple instance learning baselines for both binary and multiclass prediction tasks, enabling depth-resolved risk profiling for 3D triage based on morphological prototypes that are interpretable to pathologists.
Farfan Lopez, F. J.; Wiegering, A.; Maerkl, B.; Waidhauser, J.; Krebs, M.; Grosser, B.; Reitsam, N. G.; Probst, A.; Matthias Schrempf, M.; Schenkirsch, G.; Rosenwald, A.; Kurz, F.
Show abstract
Introduction. TAC/SARIFA has been introduced as a new robust and easy-to-evaluate biomarker in several cancer entities, including colorectal cancer. It is defined by direct contact between at least five tumour cells and one adipocyte and is believed to indicate metabolic reprogramming associated with adverse outcome. However, the mechanism that leads to TAC/SARIFA positivity remains unclear. To investigate whether there is an individual component, we conducted a study on double and triple cancers, establishing a within patient design. Methods. We retrospectively analysed a total of 135 cases with 276 colorectal cancers from two academic medical centres. The TAC/SARIFA status was evaluated, as were the basic histopathological factors. The median follow-up time was 120 months. Results. Cases with any TAC/SARIFA positive tumours showed significantly reduced overall survival (62 vs. 88 months; p = 0.011). Analysing the entire cohort, the rates of concordant and discordant cases followed a random distribution. However, restricting the analysis to synchronous pT3/4 cases revealed a significant deviation from a random distribution (p = 0.016). Conclusion. This study reveals significant concordance of TAC/SARIFA status in synchronous locally advanced colorectal double/triple carcinomas, supporting the concept that tumour adipocyte interaction reflects a host related microenvironmental condition linked to metabolic reprogramming rather than a purely tumour intrinsic event.